Papers with LLM-as-a-judge framework
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (2025.emnlp-industry)
Copied to clipboard
Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin
| Challenge: | Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions . |
| Approach: | They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples . |
| Outcome: | The proposed framework improves hallucination evaluations by leveraging human-annotated examples. |
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances (2026.eacl-long)
Copied to clipboard
| Challenge: | Common proxies such as Mean Length of Utterance (MLU), lexical diversity (vocd-D), and readability indices are dominated by length and ignore conversational context, missing aspects of response quality such as reasoning depth, topic maintenance, and discourse planning. |
| Approach: | They propose a framework that classifies the Previous Adult Utterance Type and scores the child’s response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’ s contribution to advancing the discourse). |
| Outcome: | The proposed framework assesses the child's response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’s contribution to advancing the discourse). |